feat(nucleus): Add NucleusClient.merge_model_runs() - #474
Merged
Conversation
A benchmark evaluation names a single model run, and a benchmark's items may span several datasets. A model whose predictions were uploaded as separate runs — one per dataset, or one per inference batch — therefore had no single run covering the benchmark, and every uncovered item scored as a false negative. Merging the runs produces one run that does cover it, which can then be passed to create_benchmark_evaluation_v2(). The merge is a full union: predictions are copied, never deduplicated. Colliding annotation_ids are rewritten rather than dropped, and the response reports predictions_copied, predictions_ignored and annotation_ids_rewritten so nothing is lost silently. Wraps POST /v1/nucleus/modelRun/merge. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
luke-e-schaefer
marked this pull request as ready for review
August 17, 2026 17:05
Drop the unsupported model_id parameter (the backend Joi schema rejects unknown keys and forbids cross-model merges), handle the async 202 response by returning an AsyncJob so callers wait before evaluating, make name optional to match the server default, and correct the docstring and CHANGELOG to describe the real return shape. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Contributor
|
👀 |
edwinpav
reviewed
Aug 21, 2026
edwinpav
left a comment
Contributor
There was a problem hiding this comment.
Mostly nits, one main comment here: #474 (comment)
daveguo-scale
approved these changes
Aug 21, 2026
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
edwinpav
approved these changes
Aug 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
NucleusClient.merge_model_runs(model_run_ids, name=None, *, metadata=None), wrappingPOST /v1/nucleus/modelRun/merge.Requires scaleapi PR https://github.com/scaleapi/scaleapi/pull/156919, which adds the endpoint.
Why
A benchmark evaluation names a single model run, and a benchmark's items may span several datasets. A model whose predictions were uploaded as separate runs — one per dataset, or one per inference batch — has no single run covering the benchmark, so every uncovered item scores as a false negative. Merge first, wait for the copy to finish, then pass the new run to
create_benchmark_evaluation_v2().Semantics
Asynchronous. The endpoint returns
202immediately with{"model_run_id", "dataset_ids", "job"}. The new run exists and is authorized right away, but its predictions are copied by a background job, so the run is empty until the job completes. Wait onresult["job"](anAsyncJob) before evaluating — otherwise the evaluation scores the not-yet-copied items as false negatives, the very failure this feature is meant to fix.All source runs must belong to the same model; merging across models is rejected server-side (a run's model is its provenance, read by eval, leaderboards and the model page).
Full union — predictions are copied, never deduplicated, and the source runs are left untouched. If two source runs predict on the same item with the same
annotation_id, the colliding id is rewritten rather than dropped. The copy's counts (predictions_copied,predictions_ignored,annotation_ids_rewritten) and any errors are reported on the job, not in the immediate response.Verification
The test suite requires live API keys (
conftest.pyhard-asserts onNUCLEUS_PYTEST_API_KEY), so it could not be run locally, and there is no local backend to exercise the round trip against. Verified offline against the endpoint's Joi schema and handler in scaleapi #156919: the method builds the exact payload the server accepts (model_run_ids, optionalname/metadata— nomodel_id, which the schema rejects and which the same-model constraint makes meaningless), posts tomodelRun/merge, wraps the202response'sjob_idin anAsyncJob, and raises on fewer than two distinct run ids before making a request.Note on the version bump
Bumped to 0.21.0. PR #473 also bumps
pyproject.toml/CHANGELOG.md; whichever lands second will need a trivial rebase on those two files.🤖 Generated with Claude Code
Greptile Summary
Adds
NucleusClient.merge_model_runs()for asynchronously combining predictions from multiple same-model runs into a new run.AsyncJobhandle for monitoring the background copy.Confidence Score: 5/5
The PR appears safe to merge.
No blocking failure remains.
Important Files Changed
model_run_idswire-key constant used by the merge request.Sequence Diagram
sequenceDiagram participant U as SDK caller participant C as NucleusClient participant A as Nucleus API participant J as AsyncJob participant E as Evaluation V2 U->>C: merge_model_runs(run IDs, name, metadata) C->>C: Deduplicate and validate run IDs C->>A: POST modelRun/merge A-->>C: model_run_id, dataset_ids, job metadata C-->>U: model_run_id, dataset_ids, AsyncJob U->>J: sleep_until_complete() J->>A: Poll job status A-->>J: Completed U->>E: create_benchmark_evaluation_v2(benchmark, merged run)Reviews (6): Last reviewed commit: "Merge branch 'master' into lukeschaefer/..." | Re-trigger Greptile
Context used (4)